Papers with topic discovery

6 papers
Spatial Aggregation Facilitates Discovery of Spatial Topics (P19-1)

Copied to clipboard

Challenge: a recent paper shows that topic discovery is a powerful tool for describing pertinent themes in a document corpus.
Approach: They propose to use a traditional topic discovery algorithm to find spatially distinct topics by spatial aggregation.
Outcome: The proposed method can efficiently discover spatially distinct topics from large documents.
Neural Topic Modeling with Large Language Models in the Loop (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated promising capabilities in topic discovery, but their direct application to topic modeling suffers from issues such as incomplete topic coverage, misalignment of topics, and inefficiency.
Approach: They propose a novel LLM-in-the-loop framework that integrates Large Language Models with Neural Topic Models (NTMs) global topics and document representations are learned through the NTM, while an LLM refines these topics using an Optimal Transport (OT)-based alignment objective.
Outcome: The proposed framework improves topic interpretability while preserving the efficiency of existing NTMs.
Deep Temporal-Recurrent-Replicated-Softmax for Topical Trends over Time (N18-1)

Copied to clipboard

Challenge: a novel topic model is proposed to allow topical trends to be captured in temporal collections of documents.
Approach: They propose a novel unsupervised neural dynamic topic model where topics are influenced by topic discovery over time.
Outcome: The proposed model shows better generalization, topic interpretation, evolution and trends compared to state-of-the-art models .
BlogSet-BR: A Brazilian Portuguese Blog Corpus (L18-1)

Copied to clipboard

Challenge: Several efforts have been made to build a corpus based on user-generated content . however, there is still a lack of a large semi-structured corpus that also contains author profiles in Brazilian Portuguese.
Approach: They propose to build a Brazilian Portuguese corpus with 2.1 billion words extracted from 7.4 million posts over 808 thousand different Brazilian blogs.
Outcome: The proposed corpus contains 2.1 billion words extracted from 7.4 million posts over 808 thousand different Brazilian blogs.
TAN-NTM: Topic Attention Networks for Neural Topic Modeling (2021.acl-long)

Copied to clipboard

Challenge: Topic models have been widely used to learn text representations and gain insight into document corpora.
Approach: They propose a framework which processes document as a sequence of tokens through a LSTM whose contextual outputs are attended in a topic-aware manner.
Outcome: The proposed model improves on two downstream tasks: document classification and topic guided keyphrase generation.
Siamese Network-Based Supervised Topic Modeling (D18-1)

Copied to clipboard

Challenge: Label-specific topics are widely used for supporting personality psychology, aspectlevel sentiment analysis, and crossdomain sentiment classification.
Approach: They propose a supervised topic model based on the Siamese network which trades off label-specific word distributions with document-specific label distributions in a uniform framework.
Outcome: The proposed model can trade off label-specific word distributions with document-specific label distributions in a uniform framework.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations